Abstract:To meet the practical requirements of brain-computer interfaces, the transformation of systems from traditional static design to human-machine dynamic adaptive design is regarded as one of the key methods to overcome the non-stationarity of neural signals. It contributes to the improvement of adaptive capability and overall performance of systems. Based on the fundamental theories of brain-computer interfaces and reinforcement learning, the optimization theories and applications of reinforcement learning in brain-computer interface systems are reviewed in this paper, focusing on the data acquisition, encoding/decoding and peripheral control links of non-invasive systems, as well as the relevant mechanisms for the decoding and peripheral control links of invasive systems. On the basis of the above, the online optimization mechanism of system parameters, strategies and control processes by reinforcement learning algorithms through the dynamic interaction with the environment is thoroughly investigated. The core value in enhancing the adaptability of the system to user state fluctuations, environmental noise interference and complex dynamic tasks is clarified. Finally, the current key technical challenges are analyzed and the development directions are briefly described, which provide a reference for constructing a new generation of brain-computer interface systems with long-term stability and autonomous adaptability.
[1] WOLPAW J R, BIRBAUMER N, MCFARLAND D J, et al. Brain-Computer Interfaces for Communication and Control. Clinical Neurophysiology, 2002, 113(6): 767-791. [2] CHAUDHARY U, BIRBAUMER N, RAMOS-MURGUIALDAY A.Brain-Computer Interfaces for Communication and Rehabilitation. Nature Reviews Neurology, 2016, 12(9): 513-525. [3] LOTTE F, BOUGRAIN L, CICHOCKI A, et al. A Review of Cla-ssification Algorithms for EEG-Based Brain-Computer Interfaces: A 10 Year Update. Journal of Neural Engineering, 2018, 15(3). DOI: 10.1088/1741-2552/aab2f2. [4] SHENOY P, KRAULEDAT M, BLANKERTZ B, et al. Towards Adap-tive Classification for BCI. Journal of Neural Engineering, 2006, 3(1). DOI: 10.1088/1741-2560/3/1/R02. [5] MAHMOUDI B, SANCHEZ J C.A Symbiotic Brain-Machine Interface through Value-Based Decision Making. PLoS One, 2011, 6(3). DOI: 10.1371/journal.pone.0014760. [6] PERDIKIS S, MILLAN J D R. Brain-Machine Interfaces: A Tale of Two Learners. IEEE Systems, Man, and Cybernetics Magazine, 2020, 6(3): 12-19. [7] DIGIOVANNA J, MAHMOUDI B, FORTES J, et al. Coadaptive Brain-Machine Interface via Reinforcement Learning. IEEE Transactions on Biomedical Engineering, 2009, 56(1): 54-64. [8] GIRDLER B, CALDBECK W, BAE J.Neural Decoders Using Reinforcement Learning in Brain Machine Interfaces: A Technical Review. Frontiers in Systems Neuroscience, 2022, 16. DOI: 10.3389/fnsys.2022.836778. [9] GAO X R, WANG Y J, CHEN X G, et al. Brain-Computer Interface: A Brain-in-the-Loop Communication System. Proceedings of the IEEE, 2025, 113(5): 478-511. [10] ZANDER T O, KOTHE C.Towards Passive Brain-Computer Interfaces: Applying Brain-Computer Interface Technology to Human-Machine Systems in General. Journal of Neural Engineering, 2011, 8(2). DOI: 10.1088/1741-2560/8/2/025005. [11] LEBEDEV M A, NICOLELIS M A L. Brain-Machine Interfaces: Past, Present and Future. Trends in Neurosciences, 2006, 29(9): 536-546. [12] FARWELL L A, DONCHIN E.Talking off the Top of Your Head: Toward a Mental Prosthesis Utilizing Event-Related Brain Potentials. Electroencephalography and Clinical Neurophysiology, 1988, 70(6): 510-523. [13] CHENG M, GAO X R, GAO S K, et al. Design and Implementation of a Brain-Computer Interface with High Transfer Rates. IEEE Transactions on Biomedical Engineering, 2002, 49(10): 1181-1186. [14] PFURTSCHELLER G, NEUPER C.Motor Imagery and Direct Brain-Computer Communication. Proceedings of the IEEE, 2001, 89(7): 1123-1134. [15] HILL N J, SCHÖLKOPF B. An Online Brain-Computer Interface Based on Shifting Attention to Concurrent Streams of Auditory Sti-muli. Journal of Neural Engineering, 2012, 9(2). DOI: 10.1088/1741-2560/9/2/026011. [16] PFURTSCHELLER G, SOLIS-ESCALANTE T, ORTNER R, et al. Self-Paced Operation of an SSVEP-Based Orthosis with and without an Imagery-Based “Brain Switch:” A Feasibility Study Towards a Hybrid BCI. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 2010, 18(4): 409-414. [17] NGUYEN C H, KARAVAS G K, ARTEMIADIS P.Inferring Ima-gined Speech Using EEG Signals: A New Approach Using Riema-nnian Manifold Features. Journal of Neural Engineering, 2018, 15(1). DOI: 10.1088/1741-2552/aa8235. [18] MÜLLER-GERKING J, PFURTSCHELLER G, FLYVBJERG H. Designing Optimal Spatial Filters for Single-Trial EEG Classification in a Movement Task. Clinical Neurophysiology, 1999, 110(5): 787-798. [19] HOTELLING H.Relations between Two Sets of Variates. Biometrika, 1936, 28(3/4): 321-377. [20] TANAKA H, KATURA T, SATO H.Task-Related Component Ana-lysis for Functional Neuroimaging and Application to Near-Infrared Spectroscopy Data. NeuroImage, 2013, 64: 308-327. [21] XU M P, XIAO X L, WANG Y J, et al. A Brain-Computer Interface Based on Miniature-Event-Related Potentials Induced by Very Small Lateral Visual Stimuli. IEEE Transactions on Biomedical Engineering, 2018, 65(5): 1166-1175. [22] RIVET B, SOULOUMIAC A, ATTINA V, et al. xDAWN Algorithm to Enhance Evoked Potentials: Application to Brain-Compu-ter Interface. IEEE Transactions on Biomedical Engineering, 2009, 56(8): 2035-2043. [23] BARACHANT A, BONNET S, CONGEDO M, et al. Multiclass Brain-Computer Interface Classification by Riemannian Geometry. IEEE Transactions on Biomedical Engineering, 2012, 59(4): 920-928. [24] MULLER K R, ANDERSON C W, BIRCH G E.Linear and Nonlinear Methods for Brain-Computer Interfaces. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 2003, 11(2): 165-169. [25] LOTTE F, GUAN C J.Regularizing Common Spatial Patterns to Improve BCI Designs: Unified Theory and New Algorithms. IEEE Transactions on Biomedical Engineering, 2011, 58(2): 355-362. [26] ZHANG W, WU D R.Manifold Embedded Knowledge Transfer for Brain-Computer Interfaces. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 2020, 28(5): 1117-1127. [27] SLUTZKY M W.Brain Machine Interfaces: Powerful Tools for Cli-nical Treatment and Neuroscientific Investigations. Neuroscientist, 2019, 25(2): 139-154. [28] CHEN R, CANALES A, ANIKEEVA P.Neural Recording and Modulation Technologies. Nature Reviews Materials, 2017, 2(2). DOI: 10.1038/natrevmats.2016.93. [29] ENGEMANN D A, GRAMFORT A.Automated Model Selection in Covariance Estimation and Spatial Whitening of MEG and EEG Signals. NeuroImage, 2015, 108: 328-342. [30] KAISER J F.On a Simple Algorithm to Calculate the 'energy' of a Signal // Proc of the IEEE International Conference on Acoustics, Speech, and Signal Processing. Washington, USA: IEEE, 1990: 381-384. [31] CHESTEK C A, GILJA V, NUYUJUKIAN P, et al. Long-Term Stability of Neural Prosthetic Control Signals from Silicon Cortical Arrays in Rhesus Macaque Motor Cortex. Journal of Neural Engineering, 2011, 8(4). DOI: 10.1088/1741-2560/8/4/045005. [32] BANSAL A K, TRUCCOLO W, VARGAS-IRWIN C E, et al. Decoding 3D Reach and Grasp from Hybrid Signals in Motor and Premotor Cortices: Spikes, Multiunit Activity, and Local Field Potentials. Journal of Neurophysiology, 2012, 107(5): 1337-1355. [33] GILJA V, NUYUJUKIAN P, CHESTEK C A, et al. A High-Performance Neural Prosthesis Enabled by Control Algorithm Design. Nature Neuroscience, 2012, 15(12): 1752-1757. [34] RUMMERY G A, NIRANJAN M.On-Line Q-Learning Using Connectionist Systems[C/OL]. [2026-04-16].https://www.researchgate.net/publication/2500611_On-Line_Q-Learning_Using_Connectionist_Systems. [35] MNIH V, KAVUKCUOGLU K, SILVER D, et al. Human-Level Control through Deep Reinforcement Learning. Nature, 2015, 518(7540): 529-533. [36] SCHULMAN J, LEVINE S, MORITZ P, et al. Trust Region Policy Optimization // Proc of the 32nd International Conference on Machine Learning. San Diego, USA: JMLR, 2015: 1889-1897. [37] SCHULMAN J, WOLSKI F, DHARIWAL P, et al. Proximal Policy Optimization Algorithms[C/OL].[2026-04-16]. https://arxiv.org/pdf/1707.06347. [38] HAARNOJA T, TANG H R, ABBEEL P, et al. Reinforcement Lear-ning with Deep Energy-Based Policies // Proc of the 34th International Conference on Machine Learning. San Diego, USA: JMLR, 2017: 1352-1361. [39] ZIEBART B D, MAAS A L, BAGNELL J A, et al. Maximum Entropy Inverse Reinforcement Learning // Proc of the 23rd AAAI Conference on Artificial Intelligence. Palo Alto, USA: AAAI Press, 2008: 1433-1438. [40] HO J, ERMON S.Generative Adversarial Imitation Learning[C/OL]. [2026-04-16]. https://arxiv.org/pdf/1606.03476 [41] NG A Y, RUSSELL S J.Algorithms for Inverse Reinforcement Lear-ning // Proc of the 17th International Conference on Machine Learning. San Diego, USA: JMLR, 2000: 663-670. [42] SUKHBAATAR S, SZLAM A, FERGUS R. Learning Multiagent Communication with Backpropagation // Proc of the 30th International Conference on Neural Information Processing Systems. Cambridge, USA: MIT Press, 2016: 2252-2260. [43] LILLICRAP T P, HUNT J J, PRITZEL A, et al. Continuous Control with Deep Reinforcement Learning[C/OL]. [2026-04-16]. http://arxiv.org/pdf/1509.02971. [44] FOERSTER J N, FARQUHAR G, AFOURAS T, et al. Counterfactual Multi-agent Policy Gradients. Proceedings of the AAAI Conference on Artificial Intelligence, 2018, 32(1): 2974-2982. [45] ROBBINS H.Some Aspects of the Sequential Design of Experiments. Bulletin of the American Mathematical Society, 1952, 58: 527-535. [46] LANGFORD J, ZHANG T.The Epoch-Greedy Algorithm for Multi-armed Bandits with Side Information[C/OL]. [2026-04-16].https://proceedings.neurips.cc/paper/2007/file/4b04a686b0ad13dce35fa99fa4161c65-Paper.pdf. [47] KAKADE S M, SHALEV-SHWARTZ S, TEWARI A.Efficient Bandit Algorithms for Online Multiclass Prediction[C/OL]. [2026-04-16].https://cseweb.ucsd.edu/~kamalika/teaching/CSE291W11/feb23.pdf. [48] ORSBORN A L, DANGI S, MOORMAN H G, et al. Closed-Loop Decoder Adaptation on Intermediate Time Scales Facilitates Rapid BMI Performance Improvements Independent of Decoder Initialization Conditions. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 2012, 20(4): 468-477. [49] ITURRATE I, CHAVARRIAGA R, MONTESANO L, et al. Tea-ching Brain-Machine Interfaces as an Alternative Paradigm to Neuroprosthetics Control. Scientific Reports, 2015,5(1). DOI:10.1038/srep13893. [50] PARK J, KIM K E.A POMDP Approach to Optimizing P300 Spe-ller BCI Paradigm. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 2012, 20(4): 584-594. [51] ARVANEH M, GUAN C T, ANG K K, et al. Optimizing the Cha-nnel Selection and Classification Accuracy in EEG-Based BCI. IEEE Transactions on Biomedical Engineering, 2011, 58(6): 1865-1873. [52] JIN Y X, SHANG S H, TANG L W, et al. EEG Channel Selection Algorithm Based on Reinforcement Learning // Proc of the IEEE International Conference on Networking, Sensing and Control. Wa-shington, USA: IEEE, 2022. DOI: 10.1109/ICNSC55942.2022.10004161. [53] PONGTHANISORN G, SHIRAI A, SUGIYAMA S, et al. Combination of Reinforcement and Deep Learning for EEG Channel Optimization on Brain-Machine Interface Systems // Proc of the International Conference on Artificial Intelligence in Information and Communication. Washington, USA: IEEE, 2023: 97-102. [54] PONGTHANISORN G, CAPI G.Deep Q-Learning for Channel Optimization in MRCP BMI Systems: A Teleoperated Robot Implementation. IEEE Access, 2024, 12: 73769-73778. [55] KO W, JEON E, SUK H I.A Novel RL-Assisted Deep Learning Framework for Task-Informative Signals Selection and Classification for Spontaneous BCIs. IEEE Transactions on Industrial Informatics, 2022, 18(3): 1873-1882. [56] ZHANG W, TANG X L, WANG M Z.Attention Model of EEG Signals Based on Reinforcement Learning. Frontiers in Human Neuroscience, 2024, 18. DOI: 10.3389/fnhum.2024.1442398. [57] HIEU N Q, HOANG D T, NGUYEN D N, et al. Toward BCI-Enabled Metaverse: A Joint Learning and Resource Allocation Approach // Proc of the IEEE Global Communications Conference. Washington, USA: IEEE, 2023: 3342-3348. [58] SHIN D H, SON Y H, KIM J M, et al. MARS: Multiagent Reinforcement Learning for Spatial-Spectral and Temporal Feature Selection in EEG-Based BCI. IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2024, 54(5): 3084-3096. [59] MEREL J, CARLSON D, PANINSKI L, et al. Neuroprosthetic Decoder Training as Imitation Learning. PLoS Computational Biology, 2016, 12(5). DOI: 10.1371/journal.pcbi.1004948. [60] FINN C, ABBEEL P, LEVINE S.Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks // Proc of the 34th International Conference on Machine Learning. San Diego, USA: JMLR, 2017: 1126-1135. [61] DUAN T H, CHAUHAN M, SHAIKH M A, et al. Ultra Efficient Transfer Learning with Meta Update for Cross Subject EEG Classification[C/OL]. [2026-04-16]. http://arxiv.org/pdf/2003.06113. [62] SWAMY G.Learning with Humans in the Loop: UCB/EECS-2020-76. Berkeley, USA: University of California, 2020. [63] BRYAN M J, MARTIN S A, CHEUNG W,et al. Probabilistic Co-Adaptive Brain-Computer Interfacing. Journal of Neural Enginee-ring, 2013, 10(6). DOI: 10.1088/1741-2560/10/6/066008. [64] ZHOU W Z, WU L, GAO Y K, et al. A Dynamic Window Method Based on Reinforcement Learning for SSVEP Recognition. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 2024, 32: 2114-2123. [65] NALLANI S , RAMACHANDRAN G. RLEEGNet: Integrating Brain-Computer Interfaces with Adaptive AI for Intuitive Responsiveness and High-Accuracy Motor Imagery Classification[C/OL]. [2026-04-16]. http://arxiv.org/pdf/2402.09465. [66] FIDÊNCIO A X, GRÜN F, KLAES C, et al. Hybrid Brain-Computer Interface Using Error-Related Potential and Reinforcement Learning. Frontiers in Human Neuroscience, 2025, 19. DOI: 10.3389/fnhum.2025.1569411. [67] JABRI J, HASSANHOSSEINI S, KAMALI A, et al. Reinforcement Learning-Based Feature Selection for Improving the Perfor-mance of the Brain-Computer Interface System. Signal, Image and Video Processing, 2023, 17(4): 1383-1389. [68] SALUWADANA S M R B, SILVA T. DQN-Based Framework for EEG based Stress Detection // Proc of the 10th International Conference on Information Technology Research. Washington, USA: IEEE, 2025. DOI: 10.1109/ICITR69413.2025.11353871. [69] LAMPE T, FIEDERER L D J, VOELKER M, et al. A Brain-Computer Interface for High-Level Remote Control of an Autonomous, Reinforcement-Learning-Based Robotic System for Reaching and Grasping // Proc of the 19th International Conference on Inte-lligent User Interfaces. New York, USA: ACM, 2014: 83-88. [70] DENG X Y, YU Z L, LIN C G, et al. Self-Adaptive Shared Control with Brain State Evaluation Network for Human-Wheelchair Cooperation. Journal of Neural Engineering, 2020, 17(4). DOI: 10.1088/1741-2552/ab937e. [71] XU Z C, BI L Z, YANG Z G, et al. Brain-Controlled Operator Model-Driven Deep Reinforcement Learning for Adaptive Brain-Machine Collaborative Control. Expert Systems with Applications, 2026, 305. DOI: 10.1016/j.eswa.2025.130770. [72] PHANG C R, HIRATA A.Shared Autonomy Between Human Elec-troencephalography and TD3 Deep Reinforcement Learning: A Multi-agent Copilot Approach. Annals of the New York Academy of Sciences, 2025, 1546(1): 157-172. [73] LEE J Y, LEE S, MISHRA A, et al. Brain-Computer Interface Control with Artificial Intelligence Copilots. Nature Machine Intelligence, 2025, 7(9): 1510-1523. [74] GUNASEKERA B, SAXENA T, BELLAMKONDA R, et al. Intracortical Recording Interfaces: Current Challenges to Chronic Recording Function. ACS Chemical Neuroscience, 2015, 6(1): 68-83. [75] JAROSIEWICZ B, MASSE N Y, BACHER D, et al. Advantages of Closed-Loop Calibration in Intracortical Brain-Computer Interfaces for People with Tetraplegia. Journal of Neural Engineering, 2013, 10(4). DOI: 10.1088/1741-2560/10/4/046012. [76] POHLMEYER E A, MAHMOUDI B, GENG S J, et al. Using Reinforcement Learning to Provide Stable Brain-Machine Interface Control Despite Neural Input Reorganization. PLoS One, 2014, 9(1). DOI: 10.1371/journal.pone.0087253. [77] BAE J, GIRALDO L G S, POHLMEYER E A, et al. Kernel Temporal Differences for Neural Decoding. Computational Intelligence and Neuroscience, 2015, 2015. DOI: 10.1155/2015/481375. [78] ZHANG X, CHEN S H, WANG Y W.Kernel Reinforcement Lear-ning-Assisted Adaptive Decoder Facilitates Stable and Continuous Brain Control Tasks. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 2023, 31: 4125-4134. [79] GHOSH A, SHAIKH S, ZHOU B Y, et al. Low-Complexity Reinforcement Learning Decoders for Autonomous, Scalable, Neuromorphic Intra-Cortical Brain Machine Interfaces. Neuroelectronics, 2025, 2(2). DOI: 10.55092/neuroelectronics20250006. [80] ZHOU B Y, BASU A.An Energy-Efficient Spiking Neural Network with Continuous Learning for Self-Adaptive Brain-Machine Interface. Neuromorphic Computing and Engineering, 2026, 6(2). DOI: 10.1088/2634-4386/ae6728. [81] SHAIKH S, SO R, SIBINDI T, et al. Towards Autonomous Intra-Cortical Brain Machine Interfaces: Applying Bandit Algorithms for Online Reinforcement Learning // Proc of the IEEE International Symposium on Circuits and Systems. Washington, USA: IEEE, 2020. DOI: 10.1109/ISCAS45731.2020.9180906. [82] LIANG K F, KAO J C. A Reinforcement Learning Based Software Simulator for Motor Brain-Computer Interfaces[C/OL]. [2026-04-16]. https://www.biorxiv.org/content/10.1101/2024.11.25.625180v1.full.pdf. [83] CHUA K, CALANDRA R, MCALLISTER R, et al. Deep Reinforcement Learning in a Handful of Trials Using Probabilistic Dynamics Models[C/OL].[2026-04-16]. https://arxiv.org/pdf/1805.12114. [84] WANG J X, KURTH-NELSON Z, TIRUMALA D, et al. Learning to Reinforcement Learn[C/OL]. [2026-04-16]. http://arxiv.org/pdf/1611.05763. [85] BURDA Y, EDWARDS H, STORKEY A, et al. Exploration by Random Network Distillation[C/OL]. [2026-04-16]. http://arxiv.org/pdf/1810.12894. [86] CHRISTIANO P F, LEIKE J, BROWN T, et al. Deep Reinforcement Learning from Human Preferences // Proc of the 31st International Conference on Neural Information Processing Systems. Cambridge, USA: MIT Press, 2017: 4302-4310. [87] GARCÍA-MARTÍNEZ J, MAROTO-GÓMEZ M, SEGURA-BENCOMO A, et al. Bioinspired Stimulus Selection under Multisensory Overload in Social Robots Using Reinforcement Learning. Sensors, 2025, 25(19). DOI: 10.3390/s25196152. [88] ZHANG J H, ZHU L, KONG W Z, et al. Reinforcement Learning Decoding Method of Multi-user EEG Shared Information Based on Mutual Information Mechanism. IEEE Journal of Biomedical and Health Informatics, 2025, 29(9): 6588-6598.